Skip to content

Design note: cloud compute and the run-artifact store - #348

Open
renmengye wants to merge 1 commit into
mainfrom
docs/compute-cloud-note
Open

renmengye wants to merge 1 commit into
mainfrom
docs/compute-cloud-note

Conversation

@renmengye

Copy link
Copy Markdown
Member

What

A design sketch capturing tonight's brainstorm with Mengye on how Outerloop would run on third-party/cloud compute, and the storage substrate that forces. New file docs/design/compute-cloud.md, a roadmap pointer under Beyond 1.0, and a CHANGELOG entry.

Explicitly post-0.1 — not a launch item. v0.1 ships the two battle-tested backends (Slurm + LocalCompute). This note exists so the picture is ready when the cloud version comes up.

The through-line

Slurm's shared filesystem was quietly doing three jobs — kernel state, the resident↔job channel, and run output. Cloud has no cheap fleet-wide shared FS, so each gets a home that doesn't need a shared disk:

  • git in / artifact out instead of a region-pinned shared FS → region-agnostic substrate, so jobs chase cheap capacity. A shared/parallel FS, if any, is scoped inside one multi-node job in one AZ and dies with it.
  • the resident decoupled from where experiments run: per-run backend choice (one resident drives Slurm + cloud at once), cloud VM or serverless cron, "job zero" bootstrap from the adopter's laptop, outbound-only.
  • durable state as a storage-backend choice (disk → volume → bucket), never a database — state stays blob-shaped and single-writer.
  • SkyPilot-first meta-provider so one backend covers many clouds; Modal second.
  • a run-artifact store: results, reports, traces, and full measurement trajectories as artifacts keyed by run_id, with a small metric map + provenance in the record. Traces private-by-default and bench-shaped (aligned with hermes save_sample); measurement standardized + optional; no metrics DB; W&B stays an optional live view.

Notes

  • Cross-references external.md (which names the compute + storage interfaces at a high level) and scaling.md; lands the metric-taxonomy vocabulary as the measurement-record schema.
  • Docs only. pre-commit clean.
  • Not for merge tonight — it's a design note; leaving it for a review round + Mengye's read in the morning.

🤖 Generated with Claude Code

Captures a post-0.1 direction (explicitly NOT a launch item): a third-party/
cloud compute backend and the storage substrate it forces. The through-line is
that Slurm's shared filesystem was quietly doing three jobs — kernel state,
resident<->job channel, and run output — and cloud has no cheap fleet-wide
shared FS, so each gets a home that doesn't depend on a shared disk:

- git-in / artifact-out instead of a region-pinned shared filesystem, so the
  substrate is region-agnostic and jobs can chase cheap capacity; a shared/
  parallel FS, if any, is scoped inside one multi-node job in one AZ.
- the resident decoupled from where experiments run (per-run backend choice;
  cloud VM or serverless cron; "job zero" bootstrap; drives Slurm + cloud at once).
- durable state as a storage-backend choice (disk -> volume -> bucket), never a
  database — state stays blob-shaped and single-writer.
- SkyPilot-first meta-provider so one backend covers many clouds.
- a run-artifact store: results, reports, traces, and full measurement
  trajectories as artifacts keyed by run_id; a small metric map + provenance in
  the record. Traces private-by-default and bench-shaped; measurement
  standardized and optional; no metrics DB, W&B stays an optional live view.

Roadmap pointer under Beyond 1.0; cross-references external.md and scaling.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 — reviewed head fc5e6c36 — reviewer summarizer:hermes/gpt-5.6-terra over coverage+credentials+deployment+general+lifecycle+prose.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: nothing blocking — 6 advisory notes.

6 findings attached to the lines below.

Merged verdict: the documented compute protocol is inaccurate; OUTERLOOP_ROOT is the state root; and job-local traces are not durable across preemption without checkpoint artifact uploads. Three prose rewrites are included verbatim as suggestions. Rejected: no substantive findings were rejected; repeated interface and trace-durability reports were deduplicated and retained with all agreeing lens attributions.


## The seam is already the right shape

A `compute` backend is about five operations: **submit** a job (resource spec +

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The compute interface is described with operations it does not expose. [coverage+deployment+general+prose] Compute defines submit, status, pending_reason, job_partition, active_job_names, queue_snapshot, lane_load, job_id_for_name, and cancel—not fetch, poll, capacity, or placement. Rewrite it as: "A compute backend submits, reports job state and queue information, and cancels jobs; placement is chosen by the dispatch settings."

(high confidence)

Durability is a **storage-backend choice**, in increasing weight:

- **Persistent volume (the default).** A cheap VM with an attached disk; the
state root (`OUTERLOOP_HOME`) is a mounted POSIX directory, exactly as on

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The note names the checkout path as the state root. [deployment] The chain documents OUTERLOOP_ROOT as the required state root and OUTERLOOP_HOME as the checkout; rewrite "state root (OUTERLOOP_HOME)" as "state root (OUTERLOOP_ROOT)" to avoid mounting persistent state at the wrong path.

(high confidence)

archival-science claim), age out aborted/negative ones after a window or
downsample. Same policy class as the run-dir reaper, moved to the bucket.

Because a trace streams to a local file and flushes at checkpoints + session end,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Preempted jobs lose locally flushed traces. [coverage+credentials+deployment+general+lifecycle] Traces flush only to job-local disk while the object-store artifact is sent at job end; preemption can destroy every checkpointed local trace on the ephemeral job. Specify checkpoint uploads to durable artifact storage or remove the crash-safety claim.

(high confidence)

the `compute` interface and the storage interface — plus the metric schema
sketched in the metric-taxonomy work.**

The whole design turns on one observation: Slurm gave us a shared filesystem for

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggestion. The opening explanation is padded and metaphorical. [prose] Rewrite it as: "Slurm shared filesystem stores kernel state, carries messages between the resident and jobs, and stores run output; cloud deployments need a separate service for each use."

(high confidence)

- **Ephemeral per-job first** (spin up → run → tear down); a warm pool later if
cold-start latency bites — the same resident-vs-per-cadence choice we already
made on Slurm.
- **Teardown is load-bearing** for two reasons at once: you pay per second, and a

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggestion. The teardown requirement uses a metaphor instead of stating the risk. [prose] Rewrite it as: "Teardown is required because cloud instances cost money while they run and may retain credentials after the job ends."

(high confidence)


## The run-artifact store

The final scalar was always an impoverishment. A run produces a **bundle of

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggestion. The artifact-store introduction makes an unsupported grand claim. [prose] Rewrite it as: "A scalar result omits useful outputs, so each run should store results, reports, traces, and measurement trajectories as artifacts keyed by run_id."

(high confidence)

@renmengye renmengye added the outerloop:review re-request the advisory review label Sep 28, 2026

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 1 — reviewed head fc5e6c36 — reviewer hermes/gpt-5.6-terra.

terra
Advisory findings from outerloop — the code owner decides. Reply to disagree; the outerloop:no-review label opts this PR out.

Verdict: nothing blocking — 1 advisory note.

1 finding attached to the lines below.

One advisory documentation finding: the described compute seam does not match the current interface.


## The seam is already the right shape

A `compute` backend is about five operations: **submit** a job (resource spec +

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The compute interface does not provide the five operations listed here. exposes submit, status, cancellation, and queue helpers, while output is read directly from run directories and it has no fetch or capacity/placement method, so saying both current backends already implement this interface misstates the existing seam.

(high confidence)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

outerloop:review re-request the advisory review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant